Papers with OCR quality

5 papers
Meaning Variation and Data Quality in the Corpus of Founding Era American English (2025.acl-short)

Copied to clipboard

Challenge: Legal scholars are increasingly using corpus based methods for assessing historical meaning . main corpus used in legal arguments is the Corpus of Founding Era American English .
Approach: They demonstrate how NLP can be used to infer meaning change and variation using masked language models.
Outcome: The proposed method can be used to infer meaning change and variation using advanced methods.
OCR Improves Machine Translation for Low-Resource Languages (2022.findings-acl)

Copied to clipboard

Challenge: Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages.
Approach: They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts.
Outcome: The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts.
Cheap Character Noise for OCR-Robust Multilingual Embeddings (2025.findings-acl)

Copied to clipboard

Challenge: Optical character recognition (OCR) is a key component of the digitization of historical documents.
Approach: They propose a method that fine-tunes existing multilingual models using noisy texts and a contrastive loss.
Outcome: The proposed model improves on the training data of existing models using noisy texts and a contrastive loss.
A Language Modelling Approach to Quality Assessment of OCR’ed Historical Text (2022.lrec-1)

Copied to clipboard

Challenge: a language model-based approach is used to score the quality of OCR transcriptions in the British Library Newspapers corpus . a corpus of genre-adjacent texts captures the common and legal parlance of nineteenth-century London .
Approach: They propose a language model-based approach to score the quality of OCR transcriptions in the British Library Newspapers corpus parts 1 and 2 . they aim to link newspapers of crime in nineteenth-century London to the Digital Panopticon .
Outcome: The proposed approach is based on the Proceedings of the Old Bailey Online corpus, which captures the common and legal parlance of nineteenth-century London.
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), but its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content and poor OCR quality.
Approach: They propose a corruption-robust training paradigm that surpasses existing strategies for mitigating the effects of corrupted data.
Outcome: The proposed training paradigm surpasses existing strategies for mitigating the effects of corrupted data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations